Papers by Eric Le Ferrand
Learning From Failure: Data Capture in an Australian Aboriginal Community (2022.acl-long)
Copied to clipboard
| Challenge: | a prototype of a language data capture app for speakers was tested in an Aboriginal community . elicitation of word lists, phrases, etc. has been used for decades to collect data for Indigenous languages . many software tools are developed to support linguists' work . |
| Approach: | They propose to deploy an app for speakers to confirm system guesses in an approach to transcription based on word spotting. |
| Outcome: | The proposed app was tested in an Aboriginal community in australia . it was able to confirm system guesses without a transcription bottleneck . the results were compared with other apps in the community . |
Enabling Interactive Transcription in an Indigenous Community (2020.coling-main)
Copied to clipboard
| Challenge: | Existing methods for manual transcription are often in isolation from the speech community, and so we miss out on the opportunity to take advantage of the interests and skills of local people. |
| Approach: | They propose a transcription workflow which combines spoken term detection and human-in-the-loop to support speech transcription in almost-zero resource settings. |
| Outcome: | The proposed workflow is based on two endangered languages with zero-resource datasets. |
That doesn’t sound right: Evaluating speech transcription quality in field linguistics corpora (2025.acl-short)
Copied to clipboard
| Challenge: | Automated speech recognition (ASR) is a popular tool for documenting languages, but field linguists do not have the data to train robust models. |
| Approach: | They propose to use fieldwork data to identify speech transcriptions that may be unsuitable for training ASR models. |
| Outcome: | The proposed measures can be used to identify transcriptions with characteristics common in field data but could be detrimental to ASR training. |
Are modern neural ASR architectures robust for polysynthetic languages? (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Traditional morphological typology recognizes a range of morphology in the world's languages. |
| Approach: | They investigate the performance of modern automatic speech recognition architectures on morphologically complex languages. |
| Outcome: | The proposed architectures perform better on morphologically complex languages, the authors show . they show that they are less robust in managing high OOV rates for morphology complex languages . |
How Important is a Language Model for Low-resource ASR? (2024.findings-acl)
Copied to clipboard
| Challenge: | Using an n-gram language model in ASR may seem obvious, but its absence in most implementations suggests otherwise. |
| Approach: | They examine whether using an n-gram language model in ASR can improve accuracy in low-resource languages. |
| Outcome: | The proposed model is absent in most implementations, but it does improve accuracy in English and Mandarin. |